Back

JMIR Medical Informatics

JMIR Publications Inc.

Preprints posted in the last 90 days, ranked by how well they match JMIR Medical Informatics's content profile, based on 18 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
A Human-in-the-Loop Large Language Model System Based on the Model Context Protocol for Differential Diagnosis from Electronic Medical Records and Literature

Lim, H.; Yi, H.; Yoon, J. Y.; Kwon, H.; Lee, D.; Kim, N.

2026-08-21 health informatics 10.64898/2026.08.18.26359085 medRxiv
Top 0.1%
15.0%
Show abstract

Diagnostic errors, including misdiagnoses and delayed clinical diagnoses, could affect outcomes of a significant patient population, particularly individuals presenting with rare diseases or non-specific symptoms. From rule-based diagnostic decision supporting systems (DDSS) to large language model (LLM) based tools for clinical reasoning have been developed to address these limitations. However, existing DDSS are often proprietary and difficult to integrate, and recent LLM-based tools remain hindered by operational challenges such as cost, resources constraint, and privacy concerns. Moreover, existing systems interpret electronic medical records (EMR) and generate diagnoses separately, limiting continuous evidence-based analysis and imposing repeated clinician involvement. In this paper, we present DDx-Finder, an open-source framework that leverages Model Context Protocol (MCP) servers for direct EMR and literature access, enabling prompt-driven clinical state extraction and reliable case-report re- trieval via generating searching query by LLM, while addressing limitations related to resource demands and privacy concerns. A clinical case study demonstrates the systems feasibility and its potential to provide accessible, transparent, and systematic differential diagnostic support for complex cases.

2
Global Adoption of openEHR Clinical Data Repositories: A Vendor and Community Survey

Kohler, S.; Meyer-Eschenbach, F.; Michelena, X.; Marschollek, M.; Eils, R.

2026-08-31 health informatics 10.64898/2026.08.27.26361529 medRxiv
Top 0.1%
14.7%
Show abstract

The openEHR standard provides an open, vendor-neutral architecture for clinical data repositories (CDRs), yet its real-world deployment has not been systematically documented. We conducted a dual-perspective survey combining a vendor survey of openEHR CDR providers with a community survey of openEHR practitioners. Eleven vendor organisations reported deployments across 22 countries and over 100 institutions and health regions. A complementary community survey (n=29, 17 countries) provided context on regulatory environments, adoption drivers, and barriers. Combined, the surveys cover 28 countries, 26 of them with a reported openEHR CDR deployment. Three findings emerge: openEHR has achieved national-scale presence through two distinct channels. Through vendor-market convergence, openEHR-based systems cover the majority of regional health authorities without a national mandate, including 19 of 21 Swedish regions, 3 of 4 Norwegian health regions, and 16 of 21 Finnish wellbeing services counties. Through national health record adoption, governments have built or procured national systems on openEHR as their technical foundation, including Ireland, Malta, Greece, Jamaica and Slovenia. Across Europe, this constitutes an openEHR-based interoperability infrastructure already in place across multiple EU member states. We identified no country in which openEHR is named in binding national regulation, creating structural fragility and an unrealised opportunity for alignment with the European Health Data Space (EHDS). Second, 61% of deployments serve primary use only, and 12% support both primary and secondary use. Third, lack of openEHR-specific knowledge is the most consistent adoption barrier across all geographies and deployment tiers. Adoption is driven by practitioner need and innovation, not by regulatory mandate.

3
Bridging the "Ten Walls" of Japanese Healthcare Data: A Comprehensive Semantic Mapping of JIPAD to HL7 FHIR R4 and Institutional Gap Analysis for the Japanese Health Data Space (JHDS)

Ohno, K.; Hashimoto, S.

2026-08-10 health informatics 10.64898/2026.08.06.26359847 medRxiv
Top 0.1%
12.6%
Show abstract

Background: Japan faces critical challenges in medical data interoperability, conceptualized as the "Ten Walls" obstructing the Japanese Health Data Space (JHDS) [1]. The Japanese Intensive Care Patient Database (JIPAD) - Japan's largest national ICU registry with 151 participating facilities - represents a high-quality critical care dataset that remains isolated from international data ecosystems. Objective: To develop a formal mapping of all 122 JIPAD variables to HL7 FHIR R4, characterize the nature and magnitude of semantic gaps, and assess the feasibility of JIPAD integration into the JHDS. Methods: All 122 JIPAD variables (Data Dictionary v3.7.2; Linkage Items List 20231020) were evaluated using ISO 21564 [8]-based semantic equivalence scoring across three tiers: High (direct FHIR R4 Core mapping), Partial (mapping via JP-Core Implementation Guide extensions [3]), and Low/No Equivalence (structural institutional gap). Semantically identical multi-instance fields (e.g., secondary disease codes x5) were consolidated into single mapping entries, yielding 114 mapping entries. Pseudonymization architecture was characterized from primary documentation. Results: Of 114 mapping entries representing the 122 JIPAD variables, 97 (85.1%) achieved High Equivalence via LOINC/SNOMED CT, and 12 (10.5%) achieved Partial Equivalence via JP-Core extensions, value-set translation, or FHIR R4 Core extension mechanisms - yielding a combined technical feasibility of 95.6% (109/114). Only 5 entries (4.4%) were classified as Low/No Equivalence, all attributable to Japan's proprietary disease classification system (288 adult codes; 165 pediatric codes) embedded in the DPC reimbursement framework, plus one Japan-specific procedure (PMX endotoxin adsorption) absent from international terminology systems. Variable-level mapping details are provided in Supplementary Table S1. Critically, JIPAD employs pseudonymization with record-linkage capability, enabling 99% DPC data matching - demonstrating that technical and design-level barriers to FHIR integration have already been resolved. Conclusion: JIPAD is technically and architecturally ready for FHIR integration at a 95.6% level. The remaining 4.4% barrier is exclusively institutional - rooted in MHLW policy frameworks governing the DPC disease classification system [6] - rather than technical. FHIR integration would further unlock pharmacoepidemiological and social epidemiological research currently inaccessible due to data isolation. As the sole national ICU registry providing high-acuity anchor data unavailable in general health records, JIPAD integration is essential for a clinically meaningful JHDS by 2027.

4
Adapting Clinical Event Annotation to Dutch Primary Care: An Event Annotation Framework for Post-Acute Infection Syndromes

Mazzucato, S.; Leeuwenberg, A.; van Doorn, S.; van Rosmalen, J.; Slurink, I. A. L.

2026-08-22 health informatics 10.64898/2026.08.19.26360841 medRxiv
Top 0.1%
11.7%
Show abstract

Extracting clinical information from Dutch free-text medical notes requires language-specific annotation resources, yet Dutch primary care lacks a reusable event-annotation framework for infections, post-acute infection syndromes (PAIS), and related symptoms. We adapted the COVID-19 Annotated Clinical Text (CACT) framework to Dutch and applied it to GP notes for PAIS event extraction. The framework has three annotation layers: a DiagnosticExpression typology covering acute infections, post-acute syndromes, and relevant comorbidities; an eleven-subtype Evidence inventory grounded in Dutch primary-care testing practice; and explicit decision rules for the SOEP structure of Dutch general practitioner (GP) notes (Subjective, Objective, Evaluation, Plan), including the distinction between clinician hedging and patient-side hypotheticals. On a 200-note pilot, span-level F1 under the Lybarger criterion reached 0.51 [95% CI: 0.47, 0.55] across six core entities; restricted to spans both annotators noticed, conditional F1 reached 0.78 [0.75, 0.80], indicating that most disagreement stems from annotation coverage rather than label assignment. The adaptation illustrates how an English event-based clinical annotation framework can be extended to a new language and clinical setting, yielding a reusable resource for Dutch clinical NLP; which steps generalise beyond this case (CACT to Dutch primary care) and which are specific to Dutch or PAIS remain to be tested.

5
Standardizing COVID-19 surveillance data into the OMOP common data model: a first implementation case study from Senegal

Diop, O.; Odhiambo, R.; Diouf, O.; Momanyi, R.; Ochola, M.; Diallo, A. S.; Padane, A.; Cygu, S. B.; Barasa, M.; Iddi, S.; Kiragga, A.; Sarr, M.; Mboup, S.; Mboup, A.

2026-07-06 health informatics 10.64898/2026.07.01.26357078 medRxiv
Top 0.1%
9.6%
Show abstract

The COVID-19 pandemic highlighted the need for interoperable health data infrastructures supporting reproducible observational research. The Observational Medical Outcomes Partnership Common Data Model (OMOP CDM) provides a widely adopted standard for harmonizing heterogeneous health data, but adoption remains limited in francophone Africa where language barriers and non-standardized surveillance systems pose additional challenges. We developed a complete Extract-Transform-Load (ETL) pipeline to convert a heterogeneous Senegalese COVID-19 surveillance dataset into OMOP CDM version 5.4. Source data recorded in French were translated into English through an iterative process interleaved with vocabulary mapping using ATHENA and Usagi. Semantic standardization used SNOMED CT for conditions, LOINC for measurements, and RxNorm for drugs. All 214 mappings underwent expert review by clinical and data science specialists. Data quality was assessed using the OHDSI Data Quality Dashboard (DQD) and Achilles. The standardized database achieved complete transformation (100%) for eight of the eleven source-populated domain tables, including person, visit_occurrence, measurement, and death. Partial transformation was observed for condition_occurrence (95.3%) and observation (68.1%), primarily due to incomplete vocabulary coverage for occupation categories and context-specific variables. The DQD produced an overall pass rate of 97% and a corrected pass rate of 98%, comparable to other published African OMOP implementations. Among the 19 data-quality failures, conformance and completeness issues predominated; the conformance failures were largely foreign-key checks, reflecting placeholder concept values (concept_id = 0) for metadata fields without meaningful equivalents in surveillance data. Iterative translation refinement was required when French-to-English translations did not align with OHDSI vocabulary terminology. This work documents, to our knowledge, the first OMOP CDM implementation on COVID-19 surveillance data in Senegal and francophone West Africa and provides a reusable methodological blueprint for future OMOP deployments in the region.

6
DBToken: A Database Tokenizer for Medical Event Foundation Models

Shin, I.; McCann, K.; Marino, G.; Siam, U. T.; Li, H.; Stutz, E.; Edara, R.; Loza, A. J.

2026-08-21 health informatics 10.64898/2026.08.18.26360487 medRxiv
Top 0.1%
7.6%
Show abstract

Objectives Transformer models for electronic health records require converting clinical data into token sequences, however standardized tokenization and evaluation frameworks are lacking. We introduce DBToken, an open-source library, and bits-per-row (BPR), a metric for comparing tokenization strategies. Materials and Methods DBToken accepts Medical Event Data Standard (MEDS)-compatible input and supports multiple text, numeric, and temporal tokenization strategies. BPR extends the bits-per-byte metric used in language models to enable comparison across tokenization strategies. Results DBToken efficiently tokenized data across configurations. BPR identified the vocabulary size associated with the best clinical outcome performance and localized differences in numeric tokenization performance by token class. Discussion Optimal tokenization strategies for medical foundation models are a subject of active research. DBToken enables reproducible tokenization experiments, while BPR efficiently screens vocabulary sizes and numeric representations before downstream evaluation. Conclusion DBToken and the BPR metric provide open-source infrastructure for reproducible EHR tokenization and cross-strategy evaluation.

7
Identifying patients with a phenotype consistent with chronic postsurgical pain after hip and knee arthroplasty using robust, scalable k-medoids clustering analysis

Gillam, L.; Doleman, B.; Knaggs, R.; Williams, J.

2026-08-12 orthopedics 10.64898/2026.08.11.26360161 medRxiv
Top 0.1%
7.0%
Show abstract

Background Chronic postsurgical pain (CPSP) affects between 7-23% and 13-44% of patients after hip and knee arthroplasty, respectively. Standardised methods of pain assessment provide superior evaluation of pain, including the Oxford Joint Score Pain Subscale (OJS-PS). We aim to estimate the proportion of patients with a phenotype consistent with CPSP through a k-medoids clustering technique and identify a threshold on the OJS-PS to highlight such patients at a population level. Methods In this cross-sectional study Patient Reported Outcomes Measures data 6-months after hip and knee arthroplasty from 2017 to 2025 were examined. An adapted k-medoid clustering technique utilising subsampling, batch assignment and probabilistic consensus allocated clusters. A receiver operator characteristic analysis identified a threshold on the OJS-PS noting the lowest scoring cluster. Our categorisation was compared to self-reported severe or moderate pain; sensitivity, specificity and accuracy of this categorisation were calculated. Results We analysed 109,542 hip and 113,799 knee arthroplasty patients; three clusters were used in each analysis. After hip arthroplasty: 14.4% of patients were assigned to the cluster with the lowest median OJS-PS of 11 [IQR 8 - 13]. A threshold of 15.5 classified patients as severe or moderate pain with 60.6% sensitivity, 91.0% specificity and 85.7% accuracy. Similarly, after knee arthroplasty, 25.3% were assigned to the cluster with the lowest median OJS-PS of 14 [IQR 11 - 16]. A threshold of 18.5 on the OJS-PS had an 85.4% sensitivity, 88.4% specificity and 87.8% accuracy for classifying patients with self-reported severe or moderate pain. Conclusions This robust and scalable clustering technique on ordinal clinical data estimates the proportion of patients reporting a phenotype consistent with CPSP. On a population level the thresholds identified on the OJS-PS could aid screening for potential CPSP patients 6 months after hip and knee arthroplasties.

8
Analytical Centralization of Health Expenditure at the National Administrator of Health System Resources: Architecture, Data Quality, and Operational Performance of the ADRES Health System Analytics Platform, Colombia

Garavito Jimenez, D. A.; Bello Angulo, D. E.; Mejia Lemus, L. T.; Chipatecua, D.; Fula, D. D.; Perez-Rubiano, S.; Martinez, F. L.; Bohorquez Pinzon, J. C.

2026-06-10 public and global health 10.64898/2026.06.08.26355159 medRxiv
Top 0.1%
6.9%
Show abstract

Between 2024 and 2025, Colombia universalized the Electronic Health Invoice with embedded Individual Health Services Delivery Records (RIPS -- Registro Background Between 2024 and 2025, Colombia universalized the Electronic Health Invoice with embedded RIPS records (FEV-RIPS) as the standard for financial and clinical data exchange. ADRES -- the entity responsible for administering the resources of Colombia's General Social Security Health System -- faced the challenge of processing information from multiple heterogeneous sources generated by more than 55,000 healthcare providers. Health systems in high-income countries converge clinical-financial data in consolidated platforms; Colombia started from a fragmented architecture with incompatible historical sources, no cross-database standardization, and no centralized analytical infrastructure until 2023. Objective We describe the design, technical challenges of integrating heterogeneous data, and operational performance of the analytical infrastructure built by ADRES to centralize large-scale processing of Colombian health system information, and derive transferable lessons for health system resource administrators in Latin America facing equivalent digitalization mandates. Methods Technical-descriptive report based on operational metrics from the ADRES Azure/Databricks environment during January-November 2025. We report indicators of data volume, processing speed, computational capacity, concurrent use by functional group, and governance structure. The architecture integrates VPN connectivity with MinSalud, automated processing of multiple formats (XML, relational tables, flat files), and a medallion data lake (Bronze/Silver/Gold). Data quality challenges include structural inconsistencies across sources, coding incompatibilities (municipalities, dates, diagnoses), format heterogeneities in unstructured data, and absent technical documentation. Results The platform manages 21 catalogs, 1,183 tables, and over 110,645 million stored records, with cumulative production exceeding 1 trillion processed records. It executes queries on 100 billion records in ten seconds using clusters of up to 32 TB RAM and 4,096 vCPU. During September-October 2025, monthly query peaks reached 78,028 across eleven functional groups. Integration required Python/PySpark parsers for variable-depth XML, equivalence tables for incompatible municipality codes, cleaning routines for extreme dates used as nulls (1900-01-01, 9999-12-31), and transformation logic bridging classic RIPS and FEV-RIPS. The platform supported econometric analyses, judicial mandate responses, and public interactive dashboards. Conversational AI integration (Genie, Copilot) extends analytical access to users without SQL knowledge. Conclusions ADRES built in one year an analytical infrastructure that provides, to our knowledge, the first published documentation of the systemic technical challenges of integrating heterogeneous data sources in a middle-income social security health system. Centralizing health system information at national scale is technically feasible under public institutional constraints -- but requires solving cross-source standardization problems the implementation literature does not document with quantitative precision. The derived lessons are transferable to health system resource administrators in Latin America facing equivalent challenges.

9
A consensus diabetes core dataset for research using NHS data: outputs from a Diabetes Data Science Catalyst workshop

Young, K. G.; Banerjee, A.; Dayan, C.; Denaxas, S.; Eastwood, S. V.; Jeffery, A.; Rutter, M. K.; Sattar, N.; Valabhji, J.; Horswood, R.; Humphreys, R.; Molete, M.; Murray, K.; Rogers, P.; Veiro, D.; Ireland, H.; Walker, C.; Shields, B. M.; Pearson, E. R.; McGovern, A. P.; Dennis, J. M.

2026-08-03 endocrinology 10.64898/2026.08.03.26359232 medRxiv
Top 0.1%
6.7%
Show abstract

Aims To develop a 'core' dataset of diabetes related variables to support reproducible research using UK routinely collected health data. Methods A workshop was conducted bringing together diabetes healthcare professionals, researchers, and patient and public representatives to discuss and prioritise variables for inclusion in the Diabetes Core Dataset. Core variables were those considered to be highest priority for diabetes research and available at high quality in NHS data routinely used for research (primary care [GP] and Hospital Episode Statistics [HES] data). Candidate variables for inclusion in the Diabetes Core Dataset were from a review of existing core datasets and expert opinion. Participants scored variables anonymously based on priority for diabetes research. Results 25 variables from existing diabetes core datasets and 87 other candidate variables were considered for inclusion in the Diabetes Core Dataset. All 25 of those from existing diabetes core datasets and 5 of the 87 candidate variables met the core requirements for inclusion. In addition, 7 variables were identified as high priority but not included in the core dataset as they are not currently available in GP/HES data; these were labelled as 'future high priority' variables for diabetes research. Conclusions A new diabetes core dataset for UK EHR research has been developed using a consensus-based process. The core dataset is openly available and can be flexibly applied in UK EHR (https://healthdatagateway.org/en/tool/426), including in new NHS Research Secure Data Environment platforms, to enhance reproducible research to improve the clinical care of people with diabetes and associated conditions.

10
Uncertainty-aware extraction of clinical findings from Finnish EHRs using open large language models

Leinonen, J. V.; Knuutila, J.; Kurki, S.; Pamilo, S.; Koskinen, M.

2026-07-09 health informatics 10.64898/2026.07.07.26355248 medRxiv
Top 0.1%
6.6%
Show abstract

Objective. To evaluate whether open-weight large language models (LLMs) can accurately extract clinical findings from Finnish-language pediatric records, and whether prediction uncertainty can be used to triage cases for expert review to minimize manual work. Materials and Methods. Retrospective cohort of 97 pediatric ischaemic stroke patients (1 month - 17 years) from Helsinki University Hospital (2010 - 2023). Three open LLMs (gpt-oss-20b, DeepSeek-R1-Distill-Qwen-32B, and medgemma-27b-text-it) were prompted in English to detect four extraction targets (hemiplegia, headache, seizure, and stroke as a positive control) from each patient's full free-text record. Each combination received 15 calls (five temperatures x three repeats). Performance was benchmarked against a clinician reference (accuracy, recall, precision, F1). Shannon entropy across the 15 calls quantified within-model uncertainty; inter-model disagreement provided an ensemble signal. Patients were ranked by uncertainty for a simulated selective-review workflow. Findings were externally validated in an independent neonatal stroke cohort (n = 88). Results. Gpt-oss-20b achieved the best balance of recall (0.91 - 1.00) and precision (0.83 - 0.92), with F1 0.89 - 0.95 across non-control extraction targets. Entropy in misclassified cases was 2.4 - 3.4 times higher than in correctly classified cases. Entropy-based triage achieved complete error coverage by reviewing <10% of patients for hemiplegia (8.3%) and headache (8.2%), and 19.6% for seizure. Neonatal validation reached F1 0.95 for Apgar 1 min and binary seizure, and F1 0.87 for 4-class stroke-subtype classification. Discussion. Within-model entropy and inter-model disagreement provided complementary, calibrated signals of likely error in a non-English clinical setting. Conclusion. Open LLMs can extract clinical findings from Finnish pediatric records with accuracy comparable to published English benchmarks, and uncertainty-based triage substantially reduces required expert workload.

11
Reducing Under-Triage Risk in Large Language Model Based Clinical Triage Using UMLS-CUI Augmentation

Gokhale, R.; Kukreja, M.; Kumar, N.; Gourab, K.

2026-08-10 health informatics 10.64898/2026.08.07.26358932 medRxiv
Top 0.1%
5.6%
Show abstract

Background: Public facing large language models (LLMs) are increasingly used for health guidance, including triage recommendations. We evaluated whether augmenting LLM prompts with standardized clinical concepts from the Unified Medical Language System (UMLS) could improve the safety and robustness of clinical triage recommendations. Methods: We used a publicly available dataset comprising 60 clinician-authored clinical vignettes, each represented in 16 demographic and narrative variations, yielding 960 vignette-factor combinations. Clinical entities were extracted using a two-stage pipeline combining ClinicalBERT-based named entity recognition with rule-based identification of laboratory abnormalities. Extracted entities were mapped to UMLS Concept Unique Identifiers (CUIs). Negated concepts were excluded. A confidence-weighted CUI voting classifier was trained using empirical associations between CUIs and clinician-assigned triage categories. We compared five approaches: CUI-only classification, MedGemma 27B, MedGemma 27B augmented with CUIs, GPT-4o-mini, and GPT-4o-mini augmented with CUIs. Outcomes included overall accuracy, under-triage, over-triage, emergency-case accuracy, and sensitivity to anchoring statements. Results: CUI augmentation decreased under-triage but increased over-triage in both models tested (GPT-4o-mini and MedGemma 27B). It improved high-acuity recognition while reducing recognition of low-acuity cases. CUI augmentation had mixed effects on overall triage accuracy; accuracy increased for MedGemma 27B but decreased for GPT-4o-mini. Emergency-case accuracy improved from 73.0% to 80.7% for GPT-4o-mini and from 60.5% to 68.5% for MedGemma 27B. CUI augmentation also reduced susceptibility to anchoring statements. These findings suggest that the principal value of CUI augmentation may be shifting model behavior toward safety-oriented behavior rather than uniformly improving overall accuracy. Conclusion: Ontology-grounded prompt augmentation shifted LLM triage recommendations toward greater sensitivity to high-acuity presentations and reduced overall under-triage. These safety gains were accompanied by increased over-triage and mixed effects on overall accuracy. A hybrid architecture combining LLM-based language understanding with interpretable UMLS-derived clinical concepts may improve the safety and robustness of AI-assisted triage. Further evaluation using real-world patient communications and clinical outcomes is warranted.

12
A Guided AI Framework for Customizable and Efficient Harmonisation to the OMOP Common Data Model

Nehra, N.; Swami, R.; Dadi, D.; Mishra, R.; Sharma, U.; Verma, P.; Sen, M.; Dhruw, N. K.; Jha, A. K.

2026-08-12 bioinformatics 10.64898/2026.08.07.742453 medRxiv
Top 0.1%
5.5%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWGetting clinical data from different sources to "talk" to each other within the OMOP Common Data Model (CDM) is arguably the most tedious part of multi-center research. While this integration is essential, the transformation process is frequently a manual grind, requiring a rare overlap of deep clinical knowledge and technical expertise. In this paper, we present a framework designed to alleviate some of the burden on the researcher by automating data harmonization through two distinct steps: structural schema mapping and terminological standardization. For the structural piece, we moved away from "black box" logic in favor of a stateful workflow managed by large language models (LLMs) and directed acyclic graphs. By profiling EHR data at the source, our system generates context-aware dictionaries that offer ranked mapping suggestions alongside confidence scores. While our benchmarking showed a 97.5% agreement rate at the schema level and an 84% agreement rate at the value level when compared with human experts, the system appears most effective when treated as a "co-pilot" rather than a total replacement for human oversight. To handle value-level standardization, we implemented a hybrid search strategy that pairs the semantic depth of SapBERT embeddings with the literal precision of fuzzy string matching. By using FAISS for rapid similarity retrieval, the engine attempts to resolve messy or "noisy" clinical descriptions to standard OMOP concepts. This approach seems particularly promising for handling the non-standardized labels that often plague smaller, local datasets. Ultimately, our results suggest that this guided approach can shift the timeline for OHDSI-compliant warehousing from weeks of manual curation to a more manageable and scalable pipeline, potentially lowering the barrier to entry for smaller research teams.

13
NLP Framework for Automated Symptom Severity Staging in Heart Failure and COPD Clinical Notes Using Ontology Integration: A Study Protocol

Inyangala, J.; Mukudi, F. M.; Ojino, R.; Shisanya, M. S.

2026-06-30 health informatics 10.64898/2026.06.27.26356738 medRxiv
Top 0.1%
5.5%
Show abstract

Background: Heart failure (HF) and chronic obstructive pulmonary disease (COPD) are among the leading causes of morbidity and mortality globally, with effective management heavily dependent on accurate severity staging using the New York Heart Association (NYHA) and Global Initiative for Chronic Obstructive Lung Disease (GOLD) classification systems. However, severity information is frequently embedded within unstructured clinical narratives rather than standardized Electronic Health Record (EHR) fields, limiting automated clinical decision support, disease surveillance, and retrospective healthcare analytics. Existing Natural Language Processing (NLP) approaches primarily rely on rule-based keyword extraction or supervised deep learning methods requiring large annotated corpora, which are often unavailable in many healthcare settings. Equally, most current systems inadequately integrate clinical ontologies for semantic reasoning and explainable classification, limiting interoperability and clinical applicability. Objective: This study aims to develop and evaluate an ontology-integrated NLP framework for automated extraction and severity staging of HF and COPD symptoms from de-identified clinical notes using NYHA and GOLD classification systems. Methods: The study will employ a Design Science Research (DSR) methodology to design, implement, and evaluate a hybrid NLP framework integrating rule-based extraction, SNOMED-CT ontology reasoning, and a Bidirectional Long Short-Term Memory with Conditional Random Field (Bi-LSTM-CRF) deep learning architecture for clinical sequence labeling. Approximately 1,000 de-identified clinical notes will be sampled proportionately from publicly available repositories including MIMIC-III/IV, eICU Collaborative Research Database, AmsterdamUMCdb, and MTSamples. Clinical text preprocessing will include tokenization, lemmatization, dependency parsing, abbreviation expansion, and negation detection. Ontology-guided semantic normalization will map extracted symptom entities to standardized SNOMED-CT concepts to support severity staging. Framework performance will be evaluated using precision, recall, F1-score, Cohens Kappa, sensitivity, specificity, positive predictive value, negative predictive value, confusion matrices, and correlation analyses against confirmed diagnoses and guideline-based severity classifications. Expected Outcomes: The proposed framework is expected to automate NYHA and GOLD severity staging across heterogeneous clinical note types without reliance on manually annotated severity labels. The ontology-integrated architecture is anticipated to improve semantic consistency, interpretability, and explainability of NLP outputs while enhancing EHR analytics, retrospective clinical audit, and AI-assisted clinical decision support. Conclusion: Findings from this study may provide a scalable and transferable framework for automated severity classification in data-rich but label-poor healthcare environments.

14
CPT/HCPCS Code Recommendation from Clinical Notes: A Comparative Evaluation of AI Methods

Song, Q.; Ni, C.; Liu, W.; Li, Y.; Malin, B. A.; Yin, Z.

2026-08-31 health informatics 10.64898/2026.08.29.26361731 medRxiv
Top 0.1%
5.5%
Show abstract

Automatic coding from clinical notes has been studied extensively for International Classification of Diseases (ICD) codes, yet broad Current Procedural Terminology (CPT) and Healthcare Common Procedure Coding System (HCPCS) recommendation remains comparatively underexplored. Existing studies often focus on one specialty, a limited code vocabulary, or a single model family, leaving it unclear how different artificial intelligence (AI) paradigms perform under a common, clinically meaningful evaluation. We formulate CPT and HCPCS coding as an AI-assisted recommendation task in which a physician or professional coder reviews a short, ranked list of candidate codes supported by the clinical note. Using operative notes from Vanderbilt University Medical Center (VUMC) and discharge summaries from Medical Information Mart for Intensive Care IV (MIMIC-IV), we compare lexical retrieval, Clinical-Longformer, GPT-5.6-Sol, MedGemma-27B, and an inspectable agentic-style retrieve-and-verify system under a controlled review budget. Micro-averaged recall within a fixed number of recommendations measures whether reference codes reach the reviewable list; micro-F1 is reported only where reference labels are sufficiently complete. Zero-shot GPT-5.6-Sol achieves the highest recall within five and ten candidates: 0.717 and 0.800 on VUMC and lower-bound values of 0.689 and 0.738 on MIMIC-IV. The retrieve-and-verify system reaches 0.695 and 0.784 on VUMC and lower-bound values of 0.575 and 0.657 on MIMIC-IV, with a candidate-linked evidence window attached to each retained recommendation. Diagnostic analyses reveal distinct failure sources, including output-length underfilling, confusion among closely related codes, out-of-knowledge-base generation, and incomplete evidence support. These findings establish a systematic evaluation framework for procedure-code recommendation and identify practical requirements for future systems that are accurate, review-efficient, and grounded in clinical evidence.

15
TrialCode Agent: LLM-Assisted Clinical Code-Set Construction for Trial Emulation

Habibdoust, A.; Sajjad, A.; Hernandez, D.; Patel, K.; Song, X.

2026-08-23 health informatics 10.64898/2026.08.20.26360962 medRxiv
Top 0.1%
5.4%
Show abstract

Objective Translating free-text clinical trial criteria into computable code sets is a valuable standardization practice that is necessary for producing reproducible real-world evidence studies but requires standardized interpretation across multiple clinical vocabularies. Methods We developed TrialCode Agent, a hybrid-large language model (LLM)-terminology verification agent that generates, formats, verifies, and expands candidate codes from free-text clinical criteria. The system supports ICD-9-CM diagnoses and procedures, ICD-10-CM, ICD-10-PCS, LOINC, and RxNorm medication concepts. We compared Baseline, Hybrid biomedical retrieval-augmented generation (RAG), and terminology-guided Family expansion pipelines using Claude, GPT Qwen, and MedGemma on 40 criteria from 11 trial groups. Performance was evaluated against expert-built reference code sets using exact-code precision, recall, and F1. Results The optimal pipeline varied by model. Claude with Baseline achieved the highest performance (precision 0.755, recall 0.619, F1 0.680), followed by GPT-5.5 with Baseline (precision 0.569, recall 0.658, F1 0.610), Qwen with Hybrid biomedical RAG (precision 0.656, recall 0.470, F1 0.548), and MedGemma with Family expansion (precision 0.487, recall 0.316, F1 0.383). Hybrid biomedical RAG improved aggregate F1 only for Qwen but increased GPT-5.5 RxNorm F1 from 0.320 to 0.909. Macro-averaged results showed criterion-level gains despite lower micro-averaged aggregate performance. Family expansion increased recall across models but generally reduced precision. In staged verifier ablation, micro-F1 increased from 0.254 before verification to 0.505 after final verification and expansion. Existence/vocabulary checking removed 2,594 false-positive codes, and acceptance filtering removed 952 additional false-positive codes before controlled expansion. Conclusions Combining LLM-based clinical interpretation with deterministic terminology verification produces auditable, database-ready code sets, but retrieval and broad family expansion do not consistently improve exact-code performance. Retrieval was particularly useful for RxNorm mapping, whereas overly broad or incomplete candidate generation remained the main source of error. Deterministic verification improves code validity and query readiness but cannot replace accurate clinical interpretation.

16
A Data-Driven Framework for Generating Population-Linked Case Vignettes from Nationwide Triage Data

Seidel, A.; Steiger, E.; Schuster, J.; Kroll, L. E.

2026-06-10 health informatics 10.64898/2026.06.08.26354886 medRxiv
Top 0.1%
5.4%
Show abstract

Background: Digital decision-support tools such as triage systems and symptom checkers support millions of health-related decisions each year. Their quality and safety are commonly evaluated using textual patient cases, known as case vignettes. However, existing vignette sets written by medical experts cover only a limited spectrum of real-world patient presentations and lack population weights, which would allow extrapolating evaluation results to the underlying patient population. Objective: This study aims to develop a data-driven framework for automatically generating a human-manageable set of case vignettes from nationwide triage data that captures broad presentation diversity and links each vignette to a quantitative weight reflecting the number of underlying patient assessments. Methods: From 3.2 million triage assessments conducted over one year using structured triage software in the German medical on-call service (telephone triage and online self-triage) and at the joint contact points of the outpatient emergency care service and hospital emergency departments, we randomly sampled 50,000 cases. Triage questionnaires were converted into semantic embeddings using a German Sentence Transformer Model and grouped by agglomerative clustering. For clusters containing sufficient assessments, we generated one representative assessment using a two-phase simulated-annealing optimization. The optimization minimized the distance to the cluster centroid while maximizing the number of answered triage questions, aiming for high representativeness and information content. Each representative assessment was assigned the size of its source cluster as its sample-based weight. A similarity-based sensitivity analysis was performed to examine whether these weights were preserved in the full 1-year population. Finally, the question-answer pairs of the representative assessments were converted into structured textual case vignettes using controlled prompting of a large language model. Results: The cluster analysis yielded 514 included clusters covering 96.8% of the sampled 50,000 assessments. The generated representatives showed strong agreement with the majority treatment-urgency recommendation of their source cluster (Spearman's {rho}=0.78, p<0.001) and contained on average 4.3 more answered triage questions than the original assessments within their clusters. When weighted by cluster size, the representatives approximated the sample distributions of treatment urgency, demographics, and symptoms, although some systematic deviations remained, most notably an overrepresentation of female cases (+13.5%), patients aged 14-49 years (+8.0%), and the urgency category "As soon as possible" (+6.6%). Of 121 recorded symptoms, 101 (83.5%) were covered by the representatives; the rest each occurred in <0.5% of the sample. In a sensitivity analysis, cluster-based vignette weights were strongly correlated with similarity-based population weights (Spearman's {rho}=0.77, p<0.001), and 90.1% of assessments in the full 1-year population were matched to at least one vignette. Conclusions: We present a data-driven framework for deriving a manageable set of population-weighted case vignettes from nationwide triage data. The resulting vignettes captured broad presentation diversity, approximated key sample characteristics, and provided an explicit quantitative link to the number of underlying patient assessments. After medical expert review and refinement, the vignettes may support more population-aware evaluation and quality assurance of digital decision-support tools.

17
Evaluating Clinical Concept Extraction and Evidence-Bounded Terminology Linking: Multisite Model Comparison and Pilot Ablation Study

Chen, Y.; Popescu, M.

2026-08-24 health informatics 10.64898/2026.08.20.26360740 medRxiv
Top 0.1%
5.4%
Show abstract

Background: Clinical terminology pipelines must first extract candidate spans from narrative notes and then determine whether those spans map to existing concepts or warrant further review. Evaluation is difficult because span boundaries vary between annotators and because downstream decisions depend on the terminology evidence retrieved for each span. Objective: We evaluated clinical concept extraction, terminology linking across controlled evidence conditions, and ontology-extension triage for terms that remained unmatched after initial terminology screening. Methods: We conducted 3 complementary pilot evaluations that used distinct units of analysis and were analyzed separately. Study 1 compared 5 automated extraction pipelines and a union-merge analysis with 2 human annotation sets in 66 deidentified clinical notes from 3 health systems. Agreement was evaluated by exact string matching and BGE-large-en-v1.5 embedding matching. Study 2 evaluated 56 clinical spans, including 28 with reference Unified Medical Language System concepts and 28 adjudicated as unsuitable for ontology extension, under complete retrieval, matched-concept masking, and large language model-only inference, yielding 168 span-condition outputs. The graph retrieval pipeline used BGE-large-en-v1.5 embeddings, and the decision model was Gemma 3 27B. Study 3 applied full vector retrieval to 84 terms previously not matched in either UMLS or BioPortal. Results: In Study 1, interannotator exact-match F1 was 0.29 and embedding-match F1 was 0.75. Automated exact-match F1 scores ranged from 0.07 to 0.17; embedding-match F1 was highest for MedGemma (0.55), followed by Gemma (0.53), sci_md and SciBERT (each 0.43), and Llama 3.3 (0.32). In Study 2, complete retrieval returned a reference-matched link for 28/28 known-concept spans (100%; 95% CI, 87.9%-100%). Masking assigned POSSIBLE_CANDIDATES to all 28; large language model-only inference assigned POSSIBLE_CANDIDATES to 25/28 (89.3%) and LINKED to 3/28 (10.7%). Across the 3 evidence conditions, the same 12/28 unsuitable-extension spans were classified as NOT_MEANINGFUL (42.9%) and the same 16/28 as POSSIBLE_CANDIDATES (57.1%). In Study 3, the pipeline assigned PLAUSIBLE_EXISTING_CONCEPT to all 84 terms, none was flagged for extension, and top-candidate similarity averaged 0.914 (SD 0.027); extension status was not independently adjudicated. Conclusions: Measured extraction performance varied substantially by matching definition, whereas exact-link decisions varied with the availability of matched terminology evidence. In the follow-up sample, initial nonmatching did not establish ontology novelty: after semantic retrieval, the pipeline classified all 84 terms as plausible existing concepts and proposed none for extension. These findings support separate evaluation of extraction, retrieval, evidence-grounded linking, and extension candidacy.

18
Clinical Information Needs Among Latin American Physicians: A Multi-Country Analysis of Semantic Clinical Search

Monsalve Barrientos, K.; Villa, M. C.; Castano-Villegas, N.; Zea, J.; Velasquez, L.

2026-07-09 health informatics 10.64898/2026.06.26.26356340 medRxiv
Top 0.1%
5.3%
Show abstract

Background: Physician clinical information-seeking behavior has been studied in high-income settings but remains poorly characterized in Latin America. Objective: To describe clinical information needs, geographic variation, and temporal usage patterns among physicians using a semantic clinical search platform across Latin America. Methods: We conducted a retrospective query-log analysis of physician searches performed between May 2025 and June 2026. The dataset included 235,803 queries generated by 9,443 physicians across 20 Latin American countries. Query topics were classified using a 10-category canonical taxonomy derived from platform metadata and validated through manual review of a stratified random sample of 200 queries. Results: The physician activation rate was 84.4%. Among categorized queries, Management Plan (35.6%), Differential Diagnosis (20.6%), and Work-up and Test Selection (11.4%) accounted for 67.6% of all searches. This category hierarchy was broadly consistent across the 17 countries included in country-level analyses despite differences in cohort size and healthcare settings. Query intensity was also similar across countries, with a mean of 29.6 queries per active physician over the study period. Manual validation confirmed agreement between reviewer assessment and taxonomy assignment in approximately 93% of sampled queries. Conclusions: Clinical information seeking among Latin American physicians was dominated by management planning, diagnostic reasoning, and test selection, with broadly consistent patterns across countries. These findings provide a regional behavioral baseline for understanding physician information needs in semantic clinical search systems.

19
Does OMOP CDM Conversion Improve Cross-Country Comparability of Real-World Data? A Benchmark Study in Breast Cancer and Amyotrophic Lateral Sclerosis

Aborageh, M.; Korcinska Handest, M. R.; Bakos, I.; Rajamaki, B.; Silva, C.; Horvath-Puho, E.; Pylkkaenen, L.; Venda, C.; Lentzen, M.; Becker, C.; Fernandes, J.; Paakinaho, A.; Vo, T.; Haenisch, B.; Hartikainen, S.; Tolppanen, A.-M.; Furtado, C.; Froehlich, H.; Ehrenstein, V.

2026-07-09 health informatics 10.64898/2026.07.06.26357353 medRxiv
Top 0.1%
5.3%
Show abstract

Background: Real-world data (RWD) from different countries are increasingly used to support regulatory, health technology assessment (HTA), and population-level evidence generation. However, cross-country analyses are challenged by differences in data provenance, healthcare systems, coding practices, completeness, and clinical workflows. The Observational Medical Outcomes Partnership (OMOP) common data model (CDM) is widely used to harmonise heterogeneous RWD sources, but its ability to improve comparability of downstream epidemiological analyses relative to native source data across countries requires empirical evaluation. Methods: We examined RWD from Denmark, Finland and Portugal in their ability to capture epidemiology of female breast cancer (BC) and amyotrophic lateral sclerosis (ALS), exemplifying, respectively, a common disease with established treatment modalities and high survival and a rare fatal disease with scarce treatment options. To enable head-to-head comparison on a semantic level, data were mapped to the OMOP CDM. Data in the native format were used for comparison. In a downstream analysis, we examined disease epidemiology, patient characteristics, treatment, and survival. Results: OMOP conversion enabled a common analytical framework across countries and supported semantically aligned comparisons of key epidemiological and clinical variables. However, cross-country comparability was influenced by differences in data provenance, population coverage, coding practices, availability of clinical details, treatment capture, and healthcare-system-specific workflows. Iterative comparison with native data and external clinical evidence was necessary to identify mapping issues, assess information loss, and ensure high semantic fidelity of the converted data. Overall, OMOP-based estimates were highly consistent with native-data analyses and existing clinical expectations, but residual discrepancies reflected both source-data heterogeneity and decisions in the Extract, Transform, Load (ETL) workflow design. Conclusions: OMOP CDM conversion facilitates semantically meaningful cross-country analyses of RWD by mapping heterogeneous source data to a common structure and standardised vocabularies. However, CDM conversion does not eliminate heterogeneity in the underlying data-generating processes and cannot substitute for study-specific data quality and fitness-for-purpose assessment. Robust use of harmonised RWD for regulatory, HTA, or population-level evidence generation requires iterative benchmarking against native data, clinical expertise, and data-science expertise to support valid interpretation across countries.

20
Developing an open-source framework for LLM evaluation of patients using EHR clinical documentation; performance of LLMs relative to medical professionals

Barrett, L.; Joshi, N.; North, A. S.; Dimitrov, L.; Maughan, E. F.; Ross, T.; Pankhania, R.; Paramjothy, K.; Minty, I.; Farache-Trajano, L.; Smith, S. L.; Mason, K. A.; Bhargava, E. K.; Donnelly, C.; Fatoum, H.; Padiyar, A.; Kader, Z.; Chan, C. H. K.; Schilder, A. G.; Mehta, N.

2026-08-24 otolaryngology 10.64898/2026.08.21.26361031 medRxiv
Top 0.1%
5.2%
Show abstract

Background: Large language models (LLMs) have shown increasing capability in medical knowledge tasks, yet how they perform in extracting structured clinical information from real-world clinical documentation remains uncertain. We evaluated the performance of LLMs relative to medical professionals in extracting SNOMED-coded clinical information from openly available Ear, Nose and Throat (ENT) EHRs from MTSamples, examining both reliability and accuracy metrics. Methods: We evaluated the performance of seven LLMs (including GPT-4o, Claude 3.5, Gemini 1.5 Pro, Gemma 3 and three LLAMA variants) against annotations from fourteen medical professionals who served as both study authors and data annotators. Each annotator independently extracted seven categories of clinical information from 98 publicly available ENT clinical documents: socio-demographics, symptoms, signs, diagnoses, treatments, risk factors, and test results. Standardised medical terminology was enforced through SNOMED-CT code assignment, enabling standardised comparison through Cohen's Kappa. We employed Bayesian hierarchical modelling to test non-inferiority of medic-LLM agreement compared to medic-medic agreement, using Beta distributed likelihood functions with weakly informative priors. Non-inferiority margins of 0.05, 0.10, and 0.15 were assessed with 95% posterior probability thresholds. Results: Cohen's Kappa for inter-rater reliability was 0.752 (95% CI: 0.710 - 0.794) between medical professionals and 0.391 (95% CI: 0.362-0.420) between LLMs and medical professionals. Bayesian analysis showed medic-medic agreement (posterior mean 0.813, 95% CI: 0.755-0.860) exceeded medic-LLM agreement (0.659, 95% CI: 0.633-0.684) by 0.154 (95% CI: 0.091-0.209). Non-inferiority was rejected at all tested margins (delta = 0.05, 0.10, 0.15). Agreement varied by clinical category, with smallest differences for test results and largest for diagnoses. GPT-4o achieved 97.0% precision and 84.9% recall, with a 7.5% false positive rate. Conclusions: Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation. These findings provide evidence-based guidance for LLM deployment in clinical documentation workflows, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.